Merak: An Efficient Distributed DNN Training Framework With Automated 3D Parallelism for Giant Foundation Models

نویسندگان

چکیده

Foundation models are in the process of becoming dominant deep learning technology. Pretraining a foundation model is always time-consuming due to large scale both parameter and training dataset. Besides being computing-intensive, pretraining extremely memory- communication-intensive. These challenges make it necessary apply 3D parallelism, which integrates data pipeline tensor achieve high efficiency. However, current parallelism frameworks still encounter two issues: i) they not transparent developers, requiring manual modification parallelize training, ii) their utilization computation resources, GPU memory, network bandwidth insufficient. We propose Merak , an automated framework with resource utilization. Merak automatically deploys automatic partitioner, includes graph-sharding algorithm proxy node-based graph. also offers non-intrusive API out minimal code modification. In addition, we design high-performance parallel runtime engine that employs several techniques exploit available including shifted critical path schedule increases utilization, stage-aware recomputation makes use idle worker sub-pipelined overlaps communication computation. Experiments on 64 GPUs demonstrate Merak's capability speed up performance over state-of-the-art 1.5, 2.5, 8.3, 20 billion parameters by 1.42, 1.39, 1.43, 1.61×, respectively.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

An Efficient Curvelet Framework for Denoising Images

Wiener filter suppresses noise efficiently. However, it makes the out image blurred. Curvelet preserves the edges of natural images perfectly, but, it produces visual distortion artifacts and fuzzy edges to the restored image, especially in homogeneous regions of images. In this paper, a new image denoising framework based on Curvelet transform and wiener filter is proposed, which can stop nois...

متن کامل

An MPI-Based Python Framework for Distributed Training with Keras

We present a lightweight Python framework for distributed training of neural networks on multiple GPUs or CPUs. The framework is built on the popular Keras machine learning library. The Message Passing Interface (MPI) protocol is used to coordinate the training process, and the system is well suited for job submission at supercomputing sites. We detail the software’s features, describe its use,...

متن کامل

DNN-Train: Benchmarking and Analyzing DNN Training

We aim to build a new benchmark pool for deep neural network training and to analyze how eicient existing frameworks are in performing this training. We will provide our methodology and develop proper proiling tools to perform this analysis.

متن کامل

Scalable distributed DNN training using commodity GPU cloud computing

We introduce a new method for scaling up distributed Stochastic Gradient Descent (SGD) training of Deep Neural Networks (DNN). The method solves the well-known communication bottleneck problem that arises for data-parallel SGD because compute nodes frequently need to synchronize a replica of the model. We solve it by purposefully controlling the rate of weight-update per individual weight, whic...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

ژورنال

عنوان ژورنال: IEEE Transactions on Parallel and Distributed Systems

سال: 2023

ISSN: ['1045-9219', '1558-2183', '2161-9883']

DOI: https://doi.org/10.1109/tpds.2023.3247001